01 The Big Picture
You can't manage what you can't see. Distributed systems forced the industry to invent Distributed Tracing, RED metrics, and SLOs. AI is now the same kind of dependency — a network endpoint your product's correctness rests on — with the classic observability pathologies amplified.
Your agent, copilot, or RAG pipeline is a component calling a flaky, semi-opaque service over the network. Treat it exactly like that, and three amplifications follow:
Emergent behavior
The "service" contains a model whose behavior emerges from billions of parameters you did not write and cannot enumerate. System diagrams end at your harness boundary (doc 23); past that line, behavior can only be measured, never read.
Non-determinism
Same input, different output — every time, by design. A passing test case proves one sample, not the system. Percentage-based truth is forced on you, which changes the math (§03).
Plausible wrongness
Failures come back as fluent, well-formed, confident responses. HTTP 200, p99 fine, no exception thrown — and the answer is quietly wrong. The worst failure mode of all, because it is invisible to every classic signal.
Why traditional APM isn't enough: APM assumes success is binary. For a payment service, latency and error codes cover nearly everything. For an AI feature, success has degrees — a response can be statistically right, format-valid, and on time, and still be contextually wrong: citing the retrieved document that doesn't quite answer the question, calling the right tool with subtly wrong arguments, escalating when it should have deflected. The critical signals for an AI system live above the infrastructure layer, in semantics and quality — so we need observability planes stacked on top of tracing, not instead of it.
02 The What — Four Observability Planes
AI observability is not one tool; it is a stack of four planes, each answering a different question, each reading from the trace the previous plane emitted.
Reading the four planes as one stack: the telemetry plane records facts (tokens, spans, dollars — cheap to collect, but silent about quality). The quality plane grades samples of those recorded facts, at judge prices you must budget (§03). The semantics plane watches the distribution rather than individual cases — cheap embedding math catches a traffic shift long before quality metrics move. The governance plane turns several of those signals into obligations: the audit log is built from the same spans; the redaction check runs over the same payload snapshots; the quota enforcer reads the same cost counters. Everything joins on one trace ID. That single design decision — one trace, four projections — is what makes this an observability stack rather than four disconnected dashboards.
03 The Why — SLI/SLO Math for AI
The Golden Signals translate — with an AI accent on two of them:
| Golden Signal | Classic version | AI version |
|---|---|---|
| Latency | Request duration percentiles | TTFT and ITL percentiles separately — a request can be fast to start and slow to stream (doc 24) |
| Traffic | Requests / sec | Tokens / sec per class + requests / sec — cost and load live in the token axis |
| Errors | 5xx rate, exceptions | Deflection/escalation rate · assert-based refusals · hallucination-likelihood — plus classic errors; plausible wrongness hides under a 200 |
| Saturation | CPU, queue depth | Rate-limit headroom · KV-cache pressure (doc 06, serving docs) · batch queue depth · cache-hit ratio (your doc 07 lever) |
The sampled error budget. Classic error budgets tolerate exact counts because every request is instrumented and success is observable per request. For agent tasks, "task succeeded" is often only decidable by grading — and grading costs model calls. So you grade a sample, and infer the true rate from n sampled task outcomes: p̂ = k ⁄ n failed tasks out of n total. Sampling means the number carries uncertainty, and that uncertainty must be sized — this is the same confidence math as doc 13, now online.
Wilson interval (90% CI), n = 200 sampled agent tasks, k = 14 failures → p̂ = 0.07, z = 1.645:
Two consequences: (1) with n = 200 you can wrap a 7% error rate usefully tightly — a ±3-point band — but you cannot detect a 1-point regression with this sample size; that needs n ≈ 1,000+ samples per point. (2) The error budget algebra is unchanged: budget burn = (n_failed ⁄ n_total) − (1 − SLO). If your agent-task SLO is 95% success and the sample says 93.0%, you burned 2 points of budget this period — but only know that within ±3 points, so SLA decisions on barely-breached budgets must wait for the next window or a bigger sample.
Judge cost, made explicit. The quality plane's dominant cost is the judge. Judging traffic at rate k with a judge that reads about the same tokens as the judged task costs roughly:
Numeric example: 10,000 tasks/day at ₹1.00/task = ₹10,000/day of product spend. Sampling k = 5% and judging with a small, cheap model at ₹0.25/task: 0.05 × 0.25 × 10,000 = ₹125/day — about 1.25% overhead, essentially free observability. Judge with a frontier model at ₹4/task and k = 20%, and you are paying ₹8,000/day: an 80% tax. The sampling rate and the judge price class together are one instance of the price-class thinking of doc 07 applied to monitoring.
Cost as a first-class signal. Don't just record spend — alert on it. A burn-rate alert pages when hourly token spend (per task type, per model tier) deviates beyond a few σ from a trailing baseline: a runaway agent loop that retries a failing tool twenty times shows up as spend-sigma long before any user complains. Cost telemetry is not the finance team's report; it is your saturation + runaway-failure channel in one signal.
04 How It Works — One Request, Traced End-to-End
Step through a single agent request and watch it become telemetry, then quality signals, then a dashboard update — the planes stack on one trace.
Why all seven boxes belong in one picture: each plane's data was emitted by the request itself. The cache-hit annotation is one attribute on the prefill span; the ITL percentiles are derived from the stream timestamps; the sampler scores the stored 5% of spans; the panel aggregates score samples against the SLO. One trace ID threads all four planes — no stitching, no guessing.
05 Reference Architecture — The Instrumented Pipeline
One trace per agent turn, OpenTelemetry-style spans per stage — naming follows the emerging OpenTelemetry GenAI semantic conventions conceptually; check the current spec before implementing, as of this writing the conventions are still stabilizing.
Per-stage attributes, and the alert each stage can credibly fire:
| Span name | Key attributes | Alert that could fire |
|---|---|---|
| gen_ai.completion | model, temperature, prompt_tokens {fresh, cached}, completion_tokens, ttft, itl_p95 | TTFT p99 breach · token-spend σ · cache-hit ratio drop (doc 07 lever) |
| gen_ai.tool | tool name, args-hash, outcome, duration, retries, cost | tool p99 latency · retry/timeout spike · permission-denied surge (doc 26 hooks) |
| retrieval | query, chunk_ids, hit_rate, overlap_score (doc 11 math) | retrieval hit-rate SLI below floor → RAG quality incident |
| eval.judge | trace_id of graded sample, judge model, score, rubric version | score distribution shift vs. calibrated band; judge–human disagreement ↑ |
| drift.monitor | cosine-sim shift, cluster drift %, n computed over (doc 11) | input-distribution drift gate — change-coupled releases flagged before faults surface |
The drift alert deserves emphasis: it is the cheapest early-warning signal you own. Quality metrics move only as fast as samples accumulate; embedding statistics over the input distribution move in near-real time, and you can compute them from already-logged artifacts at negligible cost — every alert above can name-check OpenTelemetry-style conventions without depending on their exactness.
06 The Evals-as-Monitors Mapping
Doc 13's eval stages are the same instruments listed in a different schedule. This table is the Rosetta stone between the two disciplines.
| Eval / monitor stage | Classic observability analogue | When it runs | What it protects |
|---|---|---|---|
| Unit check | Test in CI | Every change you make | The grader / rubric itself, on a single known case |
| Regression suite | Test suite + CI gate | Pre-merge / pre-deploy | Quality on a fixed set of mined past failures |
| Nightly suite | Synthetic check | Nightly batch over dataset | A slow-moving aggregate baseline — the "yesterday is fine" reference point |
| Canary prompts (hourly replay) | Canary deployment / synthetic probe | Hourly | Provider-side change without a version bump; early TTFT / loss signs; quality trend |
| Online judge (sampled k%) | Black-box probe / continuous SLO tracking | Always-on, sampled | The error budget itself (§03) — where p̂, CIs, and burn rates meet live traffic |
| Sampled human review | Calibrating a test environment vs. reality | Weekly sample, graded by people | The judge's own calibration — the JEV axis of doc 25 — against a ground-truth anchor |
| User feedback (👍/👎) | Satisfaction survey / CSAT | Continuous, biased | Semantically the coarsest signal — direction, never magnitude; treat as a trend channel |
Read the table once and the thesis becomes operational: eval discipline and monitoring discipline are one measurement practice on two schedules — same rubrics, same graders, same dataset lineage, all inherited from doc 13.
07 Engineering Takeaways
08 Mental Models
You don't have its source, you can't debug its internals, and the provider's SLA guarantees valid responses, not specific answers. Every hard-won discipline for third-party dependencies — health checks, synthetic probes, contract checks, exit plans, evidence retention — applies. The model is simply the dependency where the CI test suite is weakest, so production observation has to carry more weight.
The exact recast of doc 13's four-piece eval under an operational load factor. Tests run before launch over a fixed dataset you own; online monitors run after launch, on traffic you cannot curate, at prices you must budget. Same measurement math — sampling, CIs, calibration — different schedule and different assumptions. Keeping them in one codebase (same rubrics, same graders, same judge prompts) is why the mapping table of §06 actually works when it matters.
09 Common Misconceptions
"We have logs — we're observable." Logs ≠ traces ≠ evals. Logs give you what your harness wrote, unjoined; traces give you causality and latency structure through the turn loop (doc 23); evals give you the grade your logs can never self-supply. Without a judge reading the stored traffic at some sample rate, "observability" goes blind above exactly the layer where your AI product's risks live.
"LLM-as-judge is a metric, so its score is the truth." A judge is a noisy sensor, not an oracle. Its score is another model's estimate, with its own drift, truncation blindness, length bias, and self-preference. This is exactly why sampled human review (doc 25's calibrated-readout connection) is on the table: you calibrate the sensor against people, you do not replace the people with the sensor.
"Cost monitoring is finance's job." Cost is a first-class SLO input. Token spend is simultaneously your saturation signal (what's spending, at what rate, against whose quotas) and your runaway-failure channel (a stuck agent loop shows up as a spend-σ anomaly before any user traces it). A monitoring posture that only routes cost to finance discards the operational signal inside it.
"We'll add observability once the tooling is mature." It is cheap to start — one trace ID at the model-call boundary, a few structured logs, one sampled judge — and AI systems drift silently with every model update and document-set change. The "before" baseline cannot be reconstructed after the fact: instrument at launch, or your first quality incident doubles as the start of your data collection.